Papers with zero-shot multi-modal tasks
Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing fine-tuning approaches for compositional understanding compromise performance in zero-shot multi-modal tasks. |
| Approach: | They propose a method to enhance compositional understanding in pre-trained vision and language models without sacrificing performance in zero-shot multi-modal tasks. |
| Outcome: | The proposed method achieves compositionality on par with state-of-the-art models and retains strong multi-modal capabilities. |